跳转至

AI 辅助验证中的认知迁移:框架与评估协议

文章背景与核心概要

随着人工智能工具在辅助用户验证在线信息真实性方面的应用日益普及,当前的评估方法主要集中在工具运行期间的即时表现上。然而,这种评估方式忽略了用户在脱离 AI 辅助后,其独立验证能力的长期变化。

本文提出了“认知迁移”(Epistemic Transfer)这一核心概念,旨在衡量用户在使用 AI 助手后,其独立验证能力的提升程度。作者构建了一个正式的评估框架,用以区分“真正的技能习得”与“对工具的过度依赖”,并提供了一套诊断协议,帮助研究者判断 AI 工具究竟是促进了用户的认知成长,还是导致了“认知去技能化”(epistemic de-skilling)。


摘要

随着 AI 工具在辅助用户验证在线声明方面变得无处不在,当前的评估方法主要关注工具处于激活状态时的性能。本文引入了“认知迁移”(Epistemic Transfer)的概念——即用户在使用 AI 助手后,其独立验证声明的能力得到改善的程度。作者提供了一个正式框架,将此概念与单纯的依赖或信任区分开来,并提供了一套诊断协议,以确定 AI 工具是促进了真正的技能习得,还是导致了“认知去技能化”。

As AI tools become ubiquitous in assisting users with verifying online claims, current evaluation methods focus primarily on performance while the tool is active. This paper introduces the concept of Epistemic Transfer—the degree to which a user’s ability to verify claims improves independently after using an AI assistant. The author provides a formal framework to distinguish this from mere reliance or trust, offering a diagnostic protocol to determine whether AI tools foster genuine skill acquisition or lead to "epistemic de-skilling."


核心贡献

1. 定义认知迁移

作者将“认知迁移”与以下相关现象进行了区分: * 纠正效应(Correction Effects): 由于 AI 干预导致的准确性即时变化。 * 信任与依赖(Trust and Reliance): 对 AI 输出产生依赖的行为倾向。 * 人机协作性能(Human-AI Team Performance): 人类与 AI 协同工作时的综合产出。

1. Defining Epistemic Transfer

The author distinguishes "Epistemic Transfer" from related phenomena such as: * Correction Effects: Immediate changes in accuracy due to AI intervention. * Trust and Reliance: Behavioral tendencies to depend on AI outputs. * Human-AI Team Performance: The aggregate output of the human and AI working in tandem.

2. 定量指标

为了衡量 AI 辅助的长期影响,本文引入了两个主要指标: * 认知迁移效应(ETE): 通过比较不同实验条件下延迟的、无辅助的性能表现,来衡量学习收益。 * 工具移除成本(TRC): 衡量用户在撤销 AI 工具后所经历的即时性能下降程度。

2. Quantitative Metrics

To measure the long-term impact of AI assistance, the paper introduces two primary metrics: * Epistemic Transfer Effect (ETE): Compares delayed, unassisted performance across different experimental conditions to measure learning gains. * Tool-Removal Cost (TRC): Measures the immediate performance degradation experienced by a user when the AI tool is withdrawn.

3. 评估协议

所提出的协议通过整合以下要素,促进了在线实验或实地研究中的严谨测试: * 条件设置: 比较“先给答案”与“先给证据”的 AI 辅助方式。 * 对照组: 利用主动练习组和无练习对照组。 * 分析: 结合对留存声明的延迟测试,以及行为和项目层面的分析。

3. Evaluation Protocol

The proposed protocol facilitates rigorous testing in online experiments or field studies by integrating: * Conditioning: Comparing "answer-first" vs. "evidence-first" AI assistance. * Controls: Utilizing active-practice and no-practice control groups. * Analysis: Combining delayed testing on held-out claims with behavioral and item-level analysis.


诊断空间

通过映射 ETE 和 TRC 指标,该框架识别出 AI 辅助验证的四种不同结果: 1. 能力构建(Capability Building): 高 ETE;工具有效地教会了用户。 2. 能力 + 工具优势(Capability + Tool Advantage): 持续的改进与工具的持续效用相结合。 3. 认知惰性/去技能化(Epistemic Inertness / De-skilling): 低 ETE 且高 TRC;用户在没有提升自身判断力的情况下变得依赖工具。 4. 借来的验证(Verification on Loan): 工具移除后立即消失的临时性能提升。

核心论点: AI 工具的价值不应仅通过其即时辅助效果来衡量,而应通过它在用户的认知工具箱中“留下了什么”来衡量。

Diagnostic Space

By mapping the ETE and TRC metrics, the framework identifies four distinct outcomes for AI-assisted verification: 1. Capability Building: High ETE; the tool effectively teaches the user. 2. Capability + Tool Advantage: Sustained improvement combined with continued tool utility. 3. Epistemic Inertness / De-skilling: Low ETE and high TRC; the user becomes dependent on the tool without improving their own judgment. 4. Verification on Loan: Temporary performance boosts that vanish immediately upon tool removal.

Core Thesis: The value of an AI tool should not be measured solely by its immediate assistance, but by what it "leaves behind" in the user’s cognitive toolkit.


获取论文

Access the Paper